Skip to content

Optimized String bridging (Swift -> JS) - #816

Open
sliemeobn wants to merge 3 commits into
swiftwasm:mainfrom
sliemeobn:perf/fast-strings
Open

sliemeobn wants to merge 3 commits into
swiftwasm:mainfrom
sliemeobn:perf/fast-strings

Conversation

@sliemeobn

Copy link
Copy Markdown
Contributor

The basic idea is to just pass the 12 byte native representation to JS, allowing for wicked fast small string decoding and "immortal" string caching. Especially for browser use cases, where small ASCII strings are everywhere, this is quite noticeable.

In addition to a tiny bit of added complexity in the glue code, the only real downside is that we add a dependency to Swift's internal String representation, which can of course change with toolchain versions and we'd need to deal with that. I would expect that to be a somewhat infrequent event though.

I know this is a bit wild, but see a few numbers from local benchmarking below - I think it's worth it.


Apple M5 Pro · Node.js 22.21.1 · Local main (777848df) vs PR implementation. Median milliseconds per 300,000 decodes, across 30 samples. Decoder-only timings; excludes Swift lowering and Wasm calls.

Case Main Current Time reduction
ASCII, 3 bytes 13.53 ms 3.02 ms 77.7%
ASCII, 8 bytes 13.76 ms 4.00 ms 70.9%
ASCII, 9 bytes 13.88 ms 4.16 ms 70.1%
ASCII, 10 bytes 13.77 ms 4.34 ms 68.5%
Small Unicode (abc😄) 17.07 ms 7.50 ms 56.1%
Large immortal string, warm cache 14.61 ms 2.98 ms 79.6%
Large dynamic string 14.50 ms 15.01 ms −3.5%

js-framework-benchmark (in chrome on M5 Pro) with ElementaryUI (current main)
comparing JavaScriptKit 0.59.0 vs this PR

(for reference, shaving off ~10% with other means is getting really, really tough - this is the last big gun I think)

Duration

Benchmark 0.59 this PR
create rows 30.4 ms 27.5 ms
replace all rows 35.8 ms 33.3 ms
partial update 18.4 ms 18.1 ms
select row 6.8 ms 6.4 ms
swap rows 18.4 ms 18.1 ms
remove row 12.2 ms 12.5 ms
create many rows 322.3 ms 295.3 ms
append rows to large table 37.2 ms 33.7 ms
clear rows 19.2 ms 17.5 ms

Memory allocation

Benchmark 0.59 this PR
ready memory 1.07 MB 1.08 MB
run memory 4.80 MB 4.84 MB
creating/clearing 1k rows (5 cycles) 3.13 MB 3.10 MB

Transferred size and first paint

Benchmark 0.59 this PR
uncompressed size 766.7 kB 768.0 kB
compressed size (Brotli) 142.0 kB 142.3 kB
first paint 794.1 ms 818.3 ms

assert.ok(bytes.length <= 10);
const storage = new Uint8Array(12);
storage.set(bytes.subarray(0, 8));
storage[8] = bytes[8] ?? 0;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I checked this against Swift 6.4.0 wasm32 output, regular and Embedded. _SmallString.capacity is 8 on wasm32; the 10-byte capacity is for 32-bit watchOS. Real 9- and 10-byte literals take the large immortal-string path and decode correctly here.

Could we limit small() and its callers to 8 bytes, keeping the longer inputs in the large-string tests? That would keep the fixtures representative of what Swift actually emits. The decoder's cases 9/10 could then be removed too, but they do not cause a runtime failure today.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch ;)

I added these because there was discussions about re-enabling up to 10 bytes of inline strings again in Swift eventually (forgot where I read it, somewhere sunk in github PRs or issues).
So I figured this would most likely "future proof" the fast decoder.

But you are right, the 9 and 10 decoding path are currently never used as Swift does currently not produce these representations.

should we remove it?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

found the issue:
swiftlang/swift#82588

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants